scientific discovery
This Is How Anthropic Thinks AI Agents Should Navigate the Physical World
The potential for AI to automate scientific research and manufacturing must be balanced with new risks, Anthropic says. Artificial intelligence agents might occasionally get confused and hack into other computers, but Anthropic thinks it has a way to unleash the little rascals into scientific labs and manufacturing facilities safely. The AI company released details today of a new framework designed to help AI agents use physical systems like microscopes, liquid-handling equipment, quantum computing hardware, manufacturing machines, and robot arms. The framework, called Model Hardware Standard, is a set of rules that specify how AI agents should--and should not--interact with all sorts of hardware. It reflects a growing belief that AI has the potential to revolutionize scientific research and industries like manufacturing-if it can venture into the physical world safely.
How NASA is prepping for lunar traffic jams
Gateway spaceport will make Earth's air traffic control look easy. More information Adding us as a Preferred Source in Google by using this link indicates that you would like to see more of our content in Google News results. The Gateway space station will be humanity's first space station around the Moon as a vital component of the Artemis missions to return humans to the lunar surface for scientific discovery and chart the path for the first human missions to Mars. Astronauts on Gateway will be the first humans to call deep space home during missions where they will use Gateway to conduct science and prepare for lunar surface missions. Breakthroughs, discoveries, and DIY tips sent six days a week.
Interview with Akari Asai โ Beyond Scaling: Frontiers of Retrieval-Augmented Language Models
Akari describes her work on augmented language models, where language models are trained to use other models and tools. This has already resulted in OpenScholar, an open-source model which helps scientists to synthesize vast amounts of scientific literature. You were awarded the 2025 AAAI Doctoral Dissertation Award. What was the topic of your dissertation research, and why was this an interesting area of study to you? My PhD dissertation was on Retrieval-Augmented Language Models.
The Download: AI agents for science, and the "censorship-industrial complex"
Plus: A new Amazon data center could become the US's most polluting power plant. In 2024, Google DeepMind scientists shared the Nobel Prize in Chemistry for a neural network, AlphaFold, which predicts the structures of proteins. It showed that AI could make groundbreaking scientific discoveries, but AlphaFold may not be the best template for accelerating science. Instead, another approach may hold the key: AI agents. AlphaFold relied on a dataset of roughly 170,000 experimentally validated protein structures that took 53 years and roughly $21 billion worth of experimental work to assemble. Comparable datasets will be difficult or impossible to create in many fields.
The American revolutionaries who popularized science in the early United States
Benjamin Franklin and other citizen scientists are core parts of the American experiment. More information Adding us as a Preferred Source in Google by using this link indicates that you would like to see more of our content in Google News results. Benjamin Franklin's kite experiment in 1752 was a pivotal scientific event, which demonstrated the connection between lightning and electricity. Breakthroughs, discoveries, and DIY tips sent six days a week. By signing up, you confirm you are 16+, will receive newsletters and promotional content and agree to our Terms of Use and acknowledge the data practices in our Privacy Policy .
Sample Complexity of Scientific Discovery: PAC Learnability of Compositional Function Trees
Kocabay, ลuayp Talha, Akkuล, Talha Rรผzgar, Yalรงฤฑn, Kerem
Scientific discovery via symbolic regression is often viewed as statistically and computationally intractable because the hypothesis space of expressions grows combinatorially with depth. This paper revisits the statistical side through the lens of PAC learning, focusing on compositional function trees built from a finite vocabulary of smooth operators (e.g., $\{+,\times,\sin,\exp\}$ and affine maps). We prove that the relevant generalization quantity, Rademacher complexity, hence the excess risk, does not necessarily blow up exponentially with the number of distinct symbolic structures, but is controlled by (i) the depth $d$ and (ii) the Lipschitz constants of the base operators along the composed computation graph. Concretely, under mild Lipschitz conditions on operators and bounded affine leaves, a finite-union bound over a vocabulary of size $K=|\mathcal{H}_{\mathrm{base}}|$ together with Maurer-type vector contraction yields $\mathfrak{R}_n(\mathcal{H}_{\mathrm{comp}}^{d}) \leq (Kb\sqrt{2}L)^{d-1}\mathfrak{R}_n(\mathcal{H}_{\mathrm{comp}}^{1})$ with arity bound $b$; corresponding high-probability risk bounds scale as $\mathcal{O}(L^{d}/\sqrt{n})$ when $K,b=O(1)$ and $\mathfrak{R}_n(\mathcal{H}_{\mathrm{comp}}^{1})=O(n^{-1/2})$. We complement the theory with a modular codebase that trains differentiable operator trees (not MLPs) on synthetic "physics-like" targets of controlled depth and shows that the empirical generalization gap correlates positively with the predicted complexity term $(\widehat{L}^{d})/\sqrt{n}$.
Why are airplanes so cold? It's for your health.
Why are airplanes so cold? From combating fainting to helping aircraft work efficiently, planes are chilly for a reason. More information Adding us as a Preferred Source in Google by using this link indicates that you would like to see more of our content in Google News results. Breakthroughs, discoveries, and DIY tips sent six days a week. By signing up, you confirm you are 16+, will receive newsletters and promotional content and agree to our Terms of Use and acknowledge the data practices in our Privacy Policy .
Strategic Hypothesis Testing
We examine hypothesis testing within a principal-agent framework, where a strategic agent, holding private beliefs about the effectiveness of a product, submits data to a principal who decides on approval. The principal employs a hypothesis testing rule, aiming to pick a p-value threshold that balances false positives and false negatives while anticipating the agent's incentive to maximize expected profitability. Building on prior work, we develop a game-theoretic model that captures how the agent's participation and reporting behavior respond to the principal's statistical decision rule. Despite the complexity of the interaction, we show that the principal's errors exhibit clear monotonic behavior when segmented by an efficiently computable critical p-value threshold, leading to an interpretable characterization of their optimal p-value threshold.
Active Measurement: Efficient Estimation at Scale
AI has the potential to transform scientific discovery by analyzing vast datasets with little human effort. However, current workflows often do not provide the accuracy or statistical guarantees that are needed. We introduce active measurement, a human-in-the-loop AI framework for scientific measurement. An AI model is used to predict measurements for individual units, which are then sampled for human labeling using importance sampling. With each new set of human labels, the AI model is improved and an unbiased Monte Carlo estimate of the total measurement is refined. Active measurement can provide precise estimates even with an imperfect AI model, and requires little human effort when the AI model is very accurate. We derive novel estimators, weighting schemes, and confidence intervals, and show that active measurement reduces estimation error compared to alternatives in several measurement tasks.
Deciphering the Extremes: ANovel Approach for Pathological Long-tailed Recognition in Scientific Discovery
Scientific discovery across diverse fields increasingly grapples with datasets exhibiting pathological long-tailed distributions: a few common phenomena overshadow a multitude of rare yet scientifically critical instances. Unlike standard benchmarks, these scientific datasets often feature extreme imbalance coupled with a modest number of classes and limited overall sample volume, rendering existing long-tailed recognition (LTR) techniques ineffective. Such methods, biased by majority classes or prone to overfitting on scarce tail data, frequently fail to identify the very instances--novel materials, rare disease biomarkers, faint astronomical signals--that drive scientific breakthroughs. This paper introduces a novel, end-to-end framework explicitly designed to address pathological long-tailed recognition in scientific contexts. Our approach synergizes a Balanced Supervised Contrastive Learning (BSCL) mechanism, which enhances the representation of tail classes by dynamically re-weighting their contributions, with a Smooth Objective Regularization (SOR) strategy that manages the inherent tension between tail-class focus and overall classification performance. We introduce and analyze the real-world ZincFluor chemical dataset (T = 137.54)